Papers with text-only methods

6 papers
PatentVision: A multimodal method for drafting patent applications (2026.eacl-industry)

Copied to clipboard

Challenge: PatentVision integrates textual and visual inputs to generate patent specifications . existing systems fail to capture the nuanced interplay between textual, visual components .
Approach: They propose a multimodal framework that integrates textual and visual inputs to generate patent specifications.
Outcome: The proposed framework surpasses text-only methods in patent writing, the authors show . it integrates visual data to better represent intricate design features and functional connections .
Cross-media Structured Common Space for Multimedia Event Extraction (2020.acl-main)

Copied to clipboard

Challenge: We propose a new task to extract events and their arguments from multimedia documents . traditional methods target text, images or videos, but multimedia content is distributed via multimedia .
Approach: They propose a method that encodes structured representations of semantic information from textual and visual data into a common embedding space.
Outcome: The proposed method achieves 4.0% and 9.8% absolute gains on text event argument role labeling and visual event extraction.
The Truth, The Whole Truth, and Nothing but the Truth: A New Benchmark Dataset for Hebrew Text Credibility Assessment (2023.findings-emnlp)

Copied to clipboard

Challenge: a new dataset evaluates the credibility of statements made by Israeli public figures and politicians . a dataset of 1021 statements is used to assess the credibility and accuracy of statements .
Approach: They propose a dataset to evaluate the credibility of statements by Israeli politicians . they use annotated statements manually annotating them for their credibility status .
Outcome: The proposed model outperforms models based on statement and context, and achieves a 48.3 F1 score.
Bridging the Sensory Gap: Visual Injection for Taxonomy Completion (2026.acl-long)

Copied to clipboard

Challenge: Existing text-only methods suffer from a "Sensory Gap" in integrating new concepts into existing hierarchies.
Approach: They propose a framework leveraging Visual Injection for Taxonomy Completion that maps synthesized images into intrinsic pseudo-tokens and decouples magnitude from selection to prevent visual signals from being drowned out.
Outcome: Experiments on three datasets show that VITC achieves state-of-the-art performance . it delivers an average absolute gain of over 19% in Hit@1.
Point-of-Interest Type Prediction using Text and Images (2021.emnlp-main)

Copied to clipboard

Challenge: Prior efforts in POI type prediction focus on text without taking visual information into account.
Approach: They propose to use multimodal information from text and images to infer the type of a place from where a social media post was shared.
Outcome: The proposed method outperforms the state-of-the-art method for POI type prediction based on text-only methods and sheds light on cross-modal interactions and limitations.
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Multi-modal keyphrase prediction (MMKP) aims to produce concise, informative phrases that capture the essence of cross-modal inputs.
Approach: They propose to use vision-language models to generate conclusive phrases using multiple modalities of input information.
Outcome: The proposed methods outperform existing methods on absence and unseen scenarios and overestimate model capability due to overlap in training tests.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations